Back
精读 WARP-RM:通过 AR(1) 时间扭曲生成自监督进度标签,再用密集 signed progress 过滤和重加权行为克隆数据。
paper deep dive
imitation learning
reward modeling
data curation